boom layer
SHAQ: Single Headed Attention with Quasi-Recurrence
Bharwani, Nashwin, Kushner, Warren, Dandona, Sangeet, Schreiber, Ben
Natural Language Processing research has recently been dominated by large scale transformer models. Although they achieve state of the art on many important language tasks, transformers often require expensive compute resources, and days spanning to weeks to train. This is feasible for researchers at big tech companies and leading research universities, but not for scrappy start-up founders, students, and independent researchers. Stephen Merity's SHA-RNN, a compact, hybrid attention-RNN model, is designed for consumer-grade modeling as it requires significantly fewer parameters and less training time to reach near state of the art results. We analyze Merity's model here through an exploratory model analysis over several units of the architecture considering both training time and overall quality in our assessment. Ultimately, we combine these findings into a new architecture which we call SHAQ: Single Headed Attention Quasi-recurrent Neural Network. With our new architecture we achieved similar accuracy results as the SHA-RNN while accomplishing a 4x speed boost in training.
A Beginner's Guide To Attention And Memory In Deep Learning
It might have never occurred to you how you could make sense of what your friend is blabbering at a loud party. There are all kinds of noises in a party; then how come we are perfectly able to carry out a conversation? This question is known widely as the'cocktail party problem'. Most of our cognitive processes can pay attention to only a single activity at a time. In the case of a party house, our capability of directing attention towards one set of words while ignoring other sets of words, which are often overpowering, is still a conundrum.
AI Researcher Smerity Rips Apart Traditions, Coins The Term 'Boooom Layer' Because Why Not
The machine learning community this week was offered a crash course on how to write a paper by Stephen Merity (fondly called Smerity), a deep learning researcher and a Harvard graduate, via his paper on recurrent neural networks. "Stop Thinking With Your Head", announced Smerity in the title of his work, preparing the readers for a fun-filled ride. In this work, the author investigates the current state of natural language processing, the models being used and other alternate approaches. In this process, he tears down the conventional methods from top to bottom, including etymology. He not only critiques the existing methods but also provides a simplistic yet effective way to train models using minimal resources.